Back

Journal of Speech, Language, and Hearing Research

American Speech Language Hearing Association

Preprints posted in the last 90 days, ranked by how well they match Journal of Speech, Language, and Hearing Research's content profile, based on 13 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Altered Speech Processing in Childhood Listening Difficulties as Revealed by Chirped Speech Event-Related Potentials

Petley, L.; Wicks, T.; Miller, L. M.; Blankenship, C.; Chatwin, J.; Bormann, B. M.; Whittle, R. S.; Moore, D. R.

2026-08-17 otolaryngology 10.64898/2026.08.13.26360392 medRxiv
Top 0.1%
12.1%
Show abstract

Objective: Impaired understanding of noisy or degraded speech is a central feature of listening difficulties (LiD), but the possible causes of these symptoms are wide-ranging. Accordingly, recent research underscores the need to study these deficits using a test battery approach. Event-related potentials are useful objective metrics for studying LiD, but probing function across the speech processing hierarchy using traditional protocols is sequential and unrealistic in clinical settings. The novel chirped speech (Cheech) method combines natural speech with acoustic chirps to overcome these limitations. This study examines its utility for profiling childhood LiD. Methods: Twenty-eight children (15 typically developing, 13 with LiD), aged 8-17 years old, listened to a 17-minute Cheech story and detected a target word within the story via button press while EEG data were collected from 53 scalp sites. Results: Cheech successfully evoked responses from the auditory brainstem response through to the brain's language centers, as reflected by the N400 effect. Unlike TD children, those with LiD demonstrated N400 effects with atypical distributions that favored frontal rather than the typical parietal sites. A trend towards a delayed and reduced amplitude Wave V was also observed. Conclusions: Hierarchical examination of speech processing using Cheech primarily implicates altered language processing as a contributing factor to LiD, with the frontal topography of the N400 effect for those with LiD potentially suggesting a greater reliance on deliberate memory retrieval during the speech perception task. Significance: LiD could arise due to auditory and/or cognitive factors. The present results demonstrate the feasibility of objective, parallel measurement across this hierarchy and point to impaired language processing as a possible mechanism.

2
Effect of Spatial Release from Masking on Listening Effort in Different Semantic Contexts

Dantanarayana, N. D.; Li, Y.; Litovsky, R. Y.; Borjigin, A.

2026-07-20 neuroscience 10.64898/2026.07.13.738303 medRxiv
Top 0.1%
11.6%
Show abstract

Humans often communicate and learn in noisy, complex listening environments. Here, we investigated the effects of spatial hearing and semantic context cues on speech intelligibility and listening effort in young adults with typical hearing. The listening task included conditions in which target speech and speech maskers were either spatially co-located or separated. Target sentences were either semantically coherent or anomalous, while the masker comprised a mixture of two coherent sentences. Results showed higher speech intelligibility in spatially separated than co-located conditions, demonstrating a robust spatial release from masking (SRM), which is consistent with prior findings. SRM did not differ between semantically coherent and anomalous sentences, indicating comparable benefits of spatial cues across semantic contexts. However, within each spatial configuration, intelligibility was higher for coherent than anomalous sentences. Listening effort, indexed by peak pupil dilation in pupillometry measurement, was reduced in spatially separated conditions, suggesting a trend toward a release from listening effort. Analysis of the timing of peak pupil dilation revealed a significantly delayed peak dilation for anomalous sentences in the co-located condition compared with coherent sentences in the separated condition, indicating increased processing demands in the absence of spatial and semantic cues. Finally, SRM was correlated with the magnitude of release from listening effort for coherent sentences, but not for anomalous sentences, suggesting that intelligibility and listening effort benefits might co-occur when contextual cues are available.

3
Planning Difficulties in Children and Adolescents with Hearing Loss across Development

Monteseirin, K.; Mendez-Couz, M.; Rivas-Fernandez, M. A.; Conejo, N. M.

2026-08-31 psychiatry and clinical psychology 10.64898/2026.08.26.26361299 medRxiv
Top 0.1%
11.1%
Show abstract

Children and adolescents with hearing loss frequently encounter reduced auditory access and delayed language development, factors that may influence the maturation of executive functions. This study examined developmental differences in planning, a core executive function, in 98 children and adolescents with hearing loss or normal hearing aged 7 to18 years using the Tower of London task. Compared to normal hearing peers, participants with hearing loss made more unnecessary moves and rule violations and initiated problem-solving more rapidly, suggesting reduced preplanning efficiency and increased impulsivity. These group differences were most pronounced in adolescents, who showed faster initiation and greater movement inefficiency than age-matched normal hearing participants. Within the hearing loss group, adolescents displayed higher accuracy and longer initiation times than children, reflecting developmental improvements despite persistent gaps relative to hearing peers. Language development age did not alter the main effects. Findings indicate that reduced early auditory and language access may contribute to differences in planning development, highlighting the need for targeted executive functions support in educational and clinical settings for youth with hearing loss.

4
Automated Detection of Motor Speech Disorders and Subtype Classification

Wang, F.; Utianski, R. L.; Barnard, L. R.; Stricker, J. L.; Clark, H. M.; Meade, G. F.; Jones, D. T.; Whitwell, J. L.; Josephs, K. A.; Duffy, J. R.; Botha, H.

2026-07-19 neurology 10.64898/2026.07.16.26358268 medRxiv
Top 0.1%
6.8%
Show abstract

Motor speech disorders (MSDs) are early markers of neurological disease, but expert perceptual analysis is rarely available outside specialized centers. Automated speech analysis offers a scalable alternative, yet prior studies have not systematically compared modeling approaches or assessed clinically relevant metrics in independent datasets. This study compared static acoustic features, articulatory informed Phonet features, and self-supervised pretrained models for binary and multi label MSD classification. We trained and evaluated models on 583 speech samples using speaker level splits. Baseline models included logistic regression and Gated Recurrent Units (GRUs) trained on eGeMAPS and MFCCs. We extracted three types of Phonet derived features and evaluated pretrained HuBERT and SSAST models in frozen, partially fine-tuned, and fully fine-tuned configurations. Binary classification distinguished MSDs from controls, while multi label classification identified six MSD subtypes. Models were assessed using validation AUC, and cut points were tested on two independent datasets. Pretrained and Phonet based models substantially outperformed static acoustic features. In binary classification, HuBERT achieved the highest AUC (0.95), while compact Phonet derived GRUs achieved comparable performance (up to 0.94). These models generalized well to independent datasets, maintaining high sensitivity (0.94) and specificity (0.97). In multi label classification, Phonet models achieved the highest macro average AUC (0.86), but threshold-based subtype performance declined on unseen data. Automated MSD detection is feasible and clinically promising. Binary classification generalized well, whereas multi label classification showed limited threshold stability across datasets.

5
Neural tracking of stressed syllables in Dutch nursery rhymes relates to vocabulary outcomes in a large, longitudinal sample

Klis, A.;Menn, K.;Cetincelik, M.;Snijders, T.;Junge, C.

2026-06-29 Developmental Biology 10.64898/2026.06.24.734253 medRxiv
Top 0.1%
6.4%
Show abstract

Speech consists of regularities at different timescales. Already during infancy, neural electrophysiological activity aligns to these rhythms. The degree to which infants exhibit neural tracking of speech can be linked to their language development. In this study, we examined how the neural tracking of sung speech develops across age, from infancy to early childhood, and across different frequency bands (i.e., at the stress, syllabic, and phonemic rates), and whether neural tracking at each frequency and age predicts childrens language outcomes. We included 2565 children of the longitudinal YOUth cohort. Children listened to Dutch sung nursery rhymes while EEG was recorded at three measurement waves. After preprocessing the data, we included 955 children at 5 months, 1048 children at 10 months, and 795 children at 2-4 years. The final sample consisted of 750 children who also completed a receptive vocabulary test at 2-4 years. Children from 5 months onwards showed significant neural tracking of stressed syllables, syllables, and phonemes, measured with speech-brain coherence (SBC). Unexpectedly, there were no developmental changes in SBC across different frequency bands from infancy to early childhood. As expected, children with larger receptive vocabularies showed increased SBC in the stressed syllable rate. These findings suggest that stronger tracking of stressed syllables is related to individual differences in language ability.

6
Evaluating Goodness of Pronunciation and Phonological Posteriors as Objective Markers of Speech Severity in Motor Speech Disorders

Wang, F.; Utianski, R. L.; Duffy, J. R.; Barnard, L. R.; Botha, H.

2026-07-16 neurology 10.64898/2026.07.14.26358076 medRxiv
Top 0.1%
6.2%
Show abstract

This study examined the extent to which goodness of pronunciation (GoP) scores and phonological posterior probabilities capture perceptual ratings of speech severity in individuals with motor speech disorders (MSD). Speech recordings of the word catastrophe were obtained from 489 participants, including 333 neurologically typical controls and 156 individuals with MSD. GoP scores were derived using traditional acoustic features and self-supervised speech representations, including WavLM and XLS-R, across multiple modeling approaches, while phonological posterior probabilities were extracted using Phonet. Model performance was evaluated using Kendall's rank correlations, regression, and receiver operating characteristic analyses against speech-language pathologists' perceptual ratings of sound distortion and intelligibility. Both GoP and phonological posterior probabilities were significantly associated with perceptual ratings. Self-supervised speech representations substantially outperformed traditional acoustic features, with WavLM-based GoP using k-nearest neighbors achieving the strongest performance. Across correlation, regression, and classification analyses, GoP consistently outperformed phonological posterior probabilities for both sound distortion and intelligibility. Age and gender had minimal influence on model-derived measures or their relationships with perceptual ratings. These findings demonstrate the value of self-supervised GoP as an objective measure of speech impairment while highlighting the complementary role of phonological posterior probabilities in characterizing articulatory aspects of motor speech disorders.

7
Interpersonal Synchronization of Brain and Body Tracks Attention and Listening Engagement

Lambrechts, L.; Accou, B.; Vanthornhout, J.; Boets, B.; Francart, T.

2026-08-12 neuroscience 10.64898/2026.08.06.743268 medRxiv
Top 0.1%
6.1%
Show abstract

PurposeSpeech perception is a fundamental part of everyday communication that relies on more than simple identification of words and sentences. Attention and listening engagement both contribute to speech perception, while representing distinct aspects of the listening experience. Attention is typically associated with cognitive focus, whereas listening engagement additionally involves cognitive and affective immersion in sound. Despite their importance, these states remain difficult to disentangle, behaviorally and physiologically. Both have been linked to interpersonal synchronization (the synchronization of biobehavioral signals across individuals), raising questions about what this synchronization actually reflects. MethodIn this study, we disentangled attention and listening engagement by independently manipulating both factors within a single experiment. Thirty participants listened to two simultaneously presented streams of meaningful speech and were instructed to focus on only one. Both attended and unattended stimuli were designed to be either engaging or non-engaging. Neural activity was recorded using EEG, while physiological responses were measured using heart rate and electrodermal activity. ResultsInterpersonal synchronization was computed from neural and bodily signals, alongside a self- report measure of listening engagement and auditory attention decoding (AAD), a neural measure of selective attention. Interpersonal synchronization of all three modalities significantly predicted listening engagement, whereas neural interpersonal synchronization was the only measure that significantly predicted attention. These findings suggest that attention is primarily driven by cognitive processes represented in the brain, while listening engagement additionally involves affective processes that are more strongly reflected in bodily responses. ConclusionsOverall, this study demonstrates that different forms of interpersonal synchronization reflect distinct dimensions of the listening experience and supports interpersonal synchronization as a potential objective marker of listening engagement.

8
Crystallized and Fluid Cognition in Adults Who Stutter

Coalson, G.; Byrd, C. T.; Richardson, E.; Gillis, C. I.; Mahometa, M. J.

2026-08-22 public and global health 10.64898/2026.08.19.26360797 medRxiv
Top 0.1%
5.7%
Show abstract

Purpose: There is a long-standing perception that individuals who stutter are less intelligent, with the disfluencies unique to stuttered speech often assumed to be the overt reflection of lower intelligence, despite no supporting evidence. The purpose of the present study was to explore the validity of this assumption by examining the cognitive abilities of adults who stutter compared to the general population using the NIH Toolbox (C) Cognition Battery (NIHTB-CB). Method: Sixty-three adults who stutter completed the NIHTB-CB, which includes seven standardized measures assessing crystallized cognition (Picture Vocabulary, Oral Reading) and fluid cognition (List Sorting Memory Test, Pattern Comparison Processing Speed Test, Flanker Inhibitory Control Test, Dimensional Change Card Sort Test, Picture Sequence Memory Test). The NIHTB-CB generates t-scores adjusted for demographic variables based on a large sample of neurotypical adults. Results: Composite scores of overall cognition for adults who stutter were not statistically equivalent, rather, their scores were higher than the general population. Higher scores were driven by crystallized cognition, with significantly higher scores on Oral Reading subscale. Conclusions: Present findings demonstrate that adults who stutter possess cognitive skills that are comparable to or potentially higher, than the general population. These results challenge the misconception that stuttering reflects diminished intelligence and offer evidence to mitigate stereotype threat.

9
Speech clarity shapes auditory attention and visual-signal coupling during multimodal sentence comprehension

Husta, C.; Seijdel, N.; Drijvers, L.

2026-07-14 neuroscience 10.64898/2026.07.13.738151 medRxiv
Top 0.1%
5.2%
Show abstract

Face-to-face communication requires listeners to attend, integrate, and weigh multiple communicative signals, including auditory speech, mouth movements, and co-speech gestures. The contribution of these signals may depend on the reliability of auditory input and the informativeness of the available signals. We utilized rapid invisible frequency tagging (RIFT) with EEG to examine how participants attend to and integrate these different signals in clear and adverse listening conditions. Participants watched videos of an actress producing clear or noise-vocoded sentences. Auditory speech was amplitude-modulated at 58Hz, while the luminance of the gesture and mouth regions was frequency-tagged at 63Hz and 65Hz. Degraded speech elicited stronger responses at the auditory tagged frequency, suggesting increased attentional gain to the auditory signal when listening was challenging. In contrast, clear speech elicited stronger responses at the gesture tagged frequency and a stronger 2Hz intermodulation response (65-63Hz), reflecting enhanced nonlinear coupling between mouth movements and gestures. Finally, in degraded speech, the informativeness of mouth movement, but not gesture, was associated with intermodulation strength, suggesting that the informativeness of mouth movements plays a greater role in multisensory interaction when listening is challenging. Our findings demonstrate that both signal reliability and informativeness shape multisensory integration during spoken language comprehension.

10
Is that clear? Robust electrophysiological measures of the effects of prior knowledge on degraded speech perception.

Synigal, S. R.; Li, W.; Serody, M. R.; Thompson, J. L.; Lalor, E. C.

2026-08-05 neuroscience 10.64898/2026.07.31.742060 medRxiv
Top 0.1%
4.7%
Show abstract

Perception and sensation are not synonymous. Rather, perception is a process whereby sensory input is organized and interpreted in a behaviorally relevant way based on memory, experience, and context. One specific framework that is commonly invoked to explain perception is that of Bayesian inference. This framework casts perception as a probabilistic process whereby imprecise sensory data are combined with prior knowledge (prior) to determine what is consciously perceived (the posterior probability), which reflects the brains best guess as to the cause(s) of the sensory data. A striking behavioral example of how prior information can influence perception is seen in studies in which degraded speech is rendered intelligible by presenting information about the speech content in advance. Neurophysiological studies of this phenomenon have primarily focused on how it affects neural indices of low-level sensory encoding. The size of any reported effects on these indices tends to be much smaller - and much less consistent - than the notably large effects on perception that come with prior information. In the present study, we recorded EEG from 27 healthy adult participants (16 female) as they listened to degraded speech clips that were preceded by matching or mismatching text. Prior knowledge in the form of matching text led to a large perceptual pop-out effect when listening to degraded speech. Analyses of the resulting EEG revealed: 1) significant but relatively weak effects of prior information on EEG measures of the linguistic encoding of speech; and 2) a very large effect of prior information on an EEG signal that resembles a well-established neural index of perceptual evidence accumulation and that was strongly related to speech intelligibility ratings across participants. These EEG signals likely relate to separate components of a Bayesian inferential process during the predictive perception of degraded speech. As such, they have implications for understanding predictive perception more broadly and for future research on perceptual disturbances in clinical populations.

11
Auditory Working Memory and Sound Segregation Ability Predict Speech-in-Noise in Adult Cochlear Implant Users

Colak, H.; Guo, X.; Benzaquen, E.; Gurusiddappa, M.; Banerjee, A.; Choi, I.; Sedley, W.; Griffiths, T. D.

2026-06-09 neuroscience 10.64898/2026.06.05.730315 medRxiv
Top 0.1%
4.2%
Show abstract

ObjectivesOutcomes following cochlear implantation vary substantially across adult recipients, and the cognitive and perceptual factors contributing to this variability are not fully understood. This poses a challenge for developing strategies to improve cochlear implant outcomes, as such approaches require a clearer understanding of the mechanisms underlying individual listening difficulties. In this study, we investigated auditory cognitive measures in cochlear implant (CI) users to further elucidate the origins of this variability. DesignThirty-seven adult cochlear implant users completed measures of auditory cognition, comprising auditory working memory (AWM) and sound segregation ability, measured using an auditory figure-ground task (AFG), as well as measures of peripheral temporal and spectral processing, comprising the temporal modulation detection threshold (TMDT) and spectral ripple discrimination threshold (SRDT). Speech perception outcomes were assessed using word-in-noise (WIN) and sentence-in-noise (SIN) tasks. Separate multiple linear regression models evaluated the unique contribution of the auditory cognition measures to WIN and SIN performance, after accounting for the peripheral measures. ResultsBoth regression models explained a substantial proportion of variance in speech-in-noise outcomes (WIN: adjusted R{superscript 2} = 0.55; SIN: adjusted R{superscript 2}=0.57, both p < 0.001). For WIN performance, AFG and AWM were significant predictors. A similar pattern was found for SIN performance, where lower AWM ability and poorer AFG segregation were linked to poorer sentence listening in noise. No significant effects of spectral ripple discrimination or temporal modulation detection were observed in either model, even though both were significantly correlated with WIN performance. ConclusionsThese findings indicate that auditory working memory and sound segregation ability are robust predictors of speech-in-noise outcomes in adult cochlear implant users, across both word- and sentence-level measures. Together, the results may help explain why speech-in-noise outcomes remain highly variable among CI users, even when basic sensory encoding abilities are taken into account. Incorporating measures of auditory working memory and fundamental sound segregation may therefore improve outcome prediction and help in developing more individualised rehabilitation strategies.

12
Amplitude Performance Subtypes in Parkinson's Disease

Mefferd, A.; Tjaden, K.; Dietrich, M.; Brown, A. E.

2026-07-13 neurology 10.64898/2026.07.08.26357552 medRxiv
Top 0.1%
4.0%
Show abstract

Purpose: The purpose of this study was to identify subgroups of talkers with Parkinsons disease (PD) with shared tongue, lip, and jaw articulatory amplitude behaviors. The study also sought to identify demographic and clinical features that can distinguish the identified kinematic subgroups. Methods: 53 talkers with PD and 54 controls participated. Articulatory amplitudes of the tongue, lip, and jaw were measured during a paragraph reading task using three-dimensional electromagnetic articulography. Amplitude performance profiles of the tongue, lip, and jaw were established for each talker with PD by referencing their performance to that of controls. These profiles were submitted to a hierarchical cluster analysis to identify kinematic-based subgroups. Amplitude performances were compared across subgroups to determine between-group patterns. Demographic and clinical features (e.g., age, sex, disease duration, selected perceptual speech characteristics, dysarthria severity) were compared across the identified kinematic subgroups. Results: Four main kinematic subgroups with differing amplitude performance profiles were identified. One subgroup exhibited normal to mildly exaggerated or mildly reduced amplitudes and was labeled preclinical subgroup (n = 16). Three subgroups exhibited pronounced amplitude reductions of either the tongue (n = 10), the tongue and lips (n = 12), or the tongue, lips, and jaw (n = 10). In addition, there were five talkers with PD whose performance profiles did not align with the identified four subgroups. Their performance was characterized by either pronounced amplitude exaggerations or mildly reduced jaw and lip amplitudes and exaggerated tongue amplitudes. None of the demographic or clinical features differed significantly between the main four subgroups. Conclusion: Findings suggest that the extent to which hypokinesia manifests within the articulatory subsystem can vary in talkers with PD. Longitudinal studies are needed to determine if these subgroups represent different stages of disease progression or distinctly different manifestations of the disease within the articulatory subsystem.

13
Measurement and comparison of acoustic space use in vocalizations of humans and close primate relatives

Bilger, H.; J. Ryan, M.; Clarke, J.

2026-06-16 animal behavior and cognition 10.64898/2026.06.14.732185 medRxiv
Top 0.1%
3.9%
Show abstract

The human larynx, compared to those of closely related primates, lies deeper in the throat and lacks vocal membranes and air sacs. These shifts are usually analyzed regarding their acoustic effects on vowel-like vocalizations, since the evolution of speech was long thought to require an expansion of vocal range driven by vocal tract modifications. However, vowels are just one type of phoneme, and speech is just one class of human utterance. To understand the evolutionary underpinnings of known shifts in human vocal morphology, a broader bioacoustic comparison is needed. Specifically, the range of sounds used in human speech must be compared to that employed in other human vocalizations and in the repertoires of extant close primate relatives. Here, we measure the acoustic-feature space occupied by human speech, non-linguistic, and musical vocalizations along with the calls of chimpanzees, bonobos, and chacma baboons. We use Mel-frequency cepstral coefficients to create an acoustic space depicting the spectro-temporal features of over 750,000 brief vocal segments sourced from published databases and other verified sources. Speech and song occupied significantly less volume in this acoustic space than human non-linguistic vocalizations. In addition, the acoustic-feature volumes of speech and song were not statistically distinct from those of non-human primates. These results suggest that speech was not enabled by an expansion of human vocal acoustic space. Anatomical shifts unique to humans may have led to an elaboration of non-linguistic utterances, but learned vocalizations use a surprisingly small fraction of this space. Our understanding of human vocal evolution will be further informed by additional systematic comparisons of the function and homology of non-speech vocalizations, along with the collection and incorporation of more complete non-human primate vocal datasets, especially from Gorilla and Orangutan.

14
Hearing Lips and Seeing Voices After Fifty Years: A Large-Scale McGurk Illusion Dataset for Audiovisual Speech Research

Wang, Z.; Li, G.; Yu, Y.; Wu, J.; Yu, Z.; Meng, Y.; Wang, S.; Dong, C.

2026-06-10 neuroscience 10.64898/2026.06.09.731046 medRxiv
Top 0.1%
3.4%
Show abstract

Efficient face-to-face communication relies on the integration of auditory speech and visual articulatory signals. Over the past five decades, the McGurk illusion has been widely used as an index of audiovisual speech integration. However, substantial variabilities in susceptibility to the illusion across participants and speakers limit its reliability as a stable measure of audiovisual integration ability. Here, we introduce the McGurk illusion dataset (MID), which, to our knowledge, is the largest publicly available McGurk stimulus dataset to date. The MID comprises auditory (N = 400), visual (N = 400), and audiovisual (N = 640) speech stimuli generated from 80 Mandarin speakers and validated through behavioral judgments across 360,900 trials. Using this dataset, we characterized the acoustic and facial articulatory properties of McGurk stimuli, replicated substantial inter-participant and inter-speaker variabilities in illusion susceptibility, and revealed the associations between variations in McGurk illusion rate and the variations in unisensory perception, audiovisual correspondences, and speakers characteristics. Furthermore, the stimulus set enabled systematic comparisons of the reliability of different McGurk illusion-based indices of audiovisual speech integration. Overall, the MID not only provides a standardized resource for investigating audiovisual speech integration and its alterations across populations, but also supports research on speaker normalization, lip-reading, and speech perception.

15
Intelligible distracting speech disrupts early auditory attention

Richardson, B. N.; Guru Adimurthy, M.; Brown, C. A.; Ihlefeld, A.; Rosen, M. J.; Shinn-Cunningham, B. G.

2026-08-24 neuroscience 10.64898/2026.08.19.745879 medRxiv
Top 0.1%
3.3%
Show abstract

Intelligible speech disrupts selective auditory attention more than an unintelligible stream. However, low-level acoustic features of intelligible speech are relatively similar to target speech, confounding results. While controlling acoustic similarity and limiting energetic masking, we examined how masker intelligibility affects behavior and electroencephalography (EEG). Normal hearing listeners detected color words within a target stream of randomly timed words while ignoring an ongoing masker. Maskers were either spoken by the same or a different talker and comprised either isochronous sequences of intelligible words or temporally scrambled versions. Scrambled maskers either lacked broadband energy changes over time (Experiment 1) or were amplitude modulated to have the same energy profiles as intelligible, isochronous maskers (Experiment 2). In both experiments, scrambled maskers yielded better performance than intelligible maskers. For intelligible maskers, performance was better for different compared to identical talkers. EEG responses paralleled behavior: target-evoked onset responses were larger for scrambled than for intelligible maskers, particularly for identical talkers. Later target recognition responses were larger for color than other target words but unaffected by masker type or talker. Even when low-level acoustic features were carefully matched, intelligible maskers impaired auditory attention and reduced target-evoked neural responses more than scrambled maskers, implicating early sensory filtering.

16
Chirped Speech (Cheech) Enables Rapid Assessment of Multi-Level Auditory Evoked Potentials During Speech-in-Noise Recognition

Chao, M.; Holloway, C. A.; Miller, L. M.; Mankel, K.

2026-08-24 neuroscience 10.64898/2026.08.19.745831 medRxiv
Top 0.1%
3.3%
Show abstract

Difficulties understanding speech in noise remain a common complaint even among listeners with normal hearing sensitivity, highlighting the need for objective, more effective measures of real-world listening. The goal of this study was to validate the use of a novel, chirped-speech (Cheech) stimulus - continuous, naturally-spoken speech fused with chirps designed to elicit robust auditory evoked potentials - to characterize relationships between speech recognition, listening effort, and auditory neural encoding. Twenty-five normal-hearing adults completed a sentence-recognition task using both original (unmodified) and Cheech-modified AzBio sentence lists in quiet, +3 dB, and -3 dB signal-to-noise ratio (SNR) conditions while neural responses from the brainstem through cortex were recorded simultaneously. Speech recognition remained near ceiling in quiet but declined with decreasing SNR for both original and Cheech stimuli. Compared with clean speech, Cheech-modified speech showed slightly poorer recognition performance as SNR decreased and somewhat higher perceived effort overall. Yet, Cheech was highly effective at evoking auditory responses from the brainstem (auditory brainstem response, ABR) through the cortex (including middle- and late-latency responses, MLR and LLR) even with <5 minutes listening time per condition. Neural responses showed reduced amplitudes and prolonged latencies as SNR decreased. In general, ABR latencies and wave I amplitudes were associated with speech-in-noise recognition performance, whereas cortical responses (MLR Na, Nb, and LLR P1) were associated with subjective workload. These findings show that Cheech-modified speech preserves intelligibility while yielding robust, multilevel neural recordings during sentence perception, offering a promising approach to examine hierarchical auditory processing under ecologically relevant speech-in-noise conditions.

17
Investigating naming error patterns after non-invasive brain stimulation and language treatment in persons with aphasia

Sydnor, M. J.; Johnson, M. A.; Lammers, B.; Murter, J. L.; Lindquist, M.; Sebastian, R.

2026-06-16 rehabilitation medicine and physical therapy 10.64898/2026.06.08.26354856 medRxiv
Top 0.1%
3.2%
Show abstract

Abstract Background: Transcranial direct current stimulation (tDCS) paired with behavioral language therapy can improve naming in persons with aphasia (PWA), yet naming errors persist. Little is known about how naming error patterns change after non-invasive brain stimulation is combined with language treatment. Aims: To examine whether right cerebellar tDCS plus computerized aphasia therapy changes the types of naming errors in people with chronic aphasia across timepoints, and to determine whether effects differ by cerebellar tDCS polarity (anode vs. cathode). Methods and Procedures: In a randomized, double-blind, sham-controlled, within-subject crossover study, we retrospectively analyzed behavioral data from 24 individuals with post-stroke aphasia. Each participant completed two 15-session intervention periods (3-5 sessions/week) with active cerebellar tDCS + computerized aphasia therapy and sham + computerized aphasia therapy, separated by a two-month washout. General linear models (GLMs) assessed longitudinal changes in six error types (semantic, phonological real word, phonological nonword, no response, mixed, unrelated) on an untrained picture naming task (Philadelphia Naming Test; PNT) and a trained task (Naming 80; N80). Additional GLMs evaluated polarity effects with 2 (Group: anode vs. cathode) x 2 (Treatment) interactions, and treatment-order effects with 2 (Group: tDCS-first vs. sham-first) x 2 (Treatment) interactions. Outcomes and Results: Active cerebellar tDCS did not significantly change error types for trained items (N80). For untrained items (PNT), active tDCS reduced several error types relative to sham, with the clearest and most durable reduction in phonological nonword errors; more moderate reductions occurred for phonological real word and unrelated errors. Mixed errors showed a marginally opposite pattern, tending to increase after tDCS and decrease after sham. Polarity analyses indicated broadly similar effects across anodal and cathodal stimulation overall, but only the anode group showed a reliable treatment effect for phonological nonword errors on the PNT. Treatment-order analyses revealed no significant order effects. Conclusions: Our results indicate a shift in naming error types, particularly after tDCS treatment for the untrained naming task (PNT). These findings may help guide the course of treatment approaches of those with aphasia and what error naming pattern types may show changes post stroke when combining non-invasive brain stimulation and computerized aphasia therapy. Clinical Trial Registration: Cerebellar Transcranial Direct Current Stimulation and Aphasia Treatment [NCT02901574] Keywords: aphasia, naming errors, non-invasive brain stimulation, cerebellar tDCS, computerized aphasia treatment

18
Neurophysiological Evidence for Reduced Use of Prior Sound Patterns to Shape Speech Processing in Autism

Lau, J. C. Y.; McHaney, J. R.; Goldman, L.; Robinshaw, K.; Mou, F.; McFarlane, K.; Chandrasekaran, B.; Losh, M.

2026-07-10 neuroscience 10.64898/2026.07.09.737536 medRxiv
Top 0.1%
3.1%
Show abstract

Reported perceptual differences in autism may arise from reduced use of prior context to shape incoming sensory input. Speech perception provides a critical test of this account because stable perception requires listeners to integrate variable acoustic signals with contextual expectations. This study examined context-dependent modulation of speech encoding in autistic and non-autistic adults using the frequency-following response (FFR), a neurophysiological measure of phase-locked auditory encoding. Participants heard English intonational pitch contours presented in repetitive and variable contexts while EEG was recorded. Principal component analysis of FFR metrics yielded components indexing neural encoding fidelity and timing. Non-autistic participants showed enhanced encoding fidelity in more predictable contexts, whereas autistic participants showed reduced context-dependent modulation. Neural encoding timing also showed divergent context effects across groups, suggesting altered balance between feedback-based predictive mechanisms and locally driven adaptation processes. Within the autistic group, greater context-related modulation of encoding fidelity was associated with lower ADOS-2 Social Affect severity but poorer speech-in-noise perception, suggesting that the functional impact of contextual modulation depends on input reliability and task demands. These findings indicate that context-dependent modulation of speech encoding is altered in autism and may contribute to individual differences in auditory and social-communicative function.

19
Age-related changes in acoustic cue use for speech-in-speech perception

Fish, E.; DiNino, M.

2026-06-22 otolaryngology 10.64898/2026.06.17.26355866 medRxiv
Top 0.1%
2.8%
Show abstract

Acoustic cues such as pitch and spatial location allow listeners to attend to a target speaker and ignore competing talkers, aiding speech recognition in background noise. Diminished ability to utilize acoustic cues for speech stream segregation may thus contribute to older adults' challenges hearing in noise. Adults aged 18-74 completed a speech-in-speech identification task with three conditions containing 1) only pitch cues (fundamental frequency), 2) only spatial cues (interaural time differences; ITDs), and 3) both pitch and spatial cues for segregating a target talker from competing talkers. Hearing thresholds at standard and extended high frequencies (EHFs), auditory brainstem responses (ABRs), and digit span scores were acquired to examine the influence of sensory and cognitive factors on use of each acoustic cue for speech-in-speech recognition. Significant differences were observed between cue condition scores indicating that use of the available cue(s) drove performance. ABR metrics were not a significant predictor but digit span scores significantly predicted scores on all three cue conditions. Working memory abilities therefore set a baseline for participants' speech-in-speech recognition regardless of the acoustic content. Hearing thresholds at standard frequencies significantly predicted scores on the Pitch condition. EHF hearing thresholds better predicted Spatial and Both Cue condition performance, suggesting that EHF thresholds represent auditory processing important for coding ITDs. Age group analysis revealed that older adults (aged 40+) performed significantly more poorly on all cue conditions of the speech-in-speech recognition task relative to younger adults. Age-related changes in auditory sensory processing may therefore impair older adults' speech-in-noise perception by reducing their ability to use acoustic cues for segregating target and competing speech.

20
The Effect of Plosive Content on the Loudness Perception of Vowel-Consonant-Vowel Syllables in Listeners with Sensorineural Hearing Loss

Davies, T.; Bleeck, S.

2026-08-27 neuroscience 10.64898/2026.08.26.747282 medRxiv
Top 0.1%
2.7%
Show abstract

Objective: This study investigated whether plosive consonants carry a perceptual loudness weighting that significantly exceeds that of non-plosive consonants when judged by hearing-impaired listeners. Design: A prospective loudness matching experiment utilizing the method of adjustment. Study Sample: 19 consenting native English speakers (Mean age: 61.4, SD: 16.4) with bilateral mild to moderate high-frequency sensorineural hearing loss, indicative of presbycusis. Stimuli: 13 vowel-consonant-vowel (VCV) nonsense syllables, exclusively utilizing the flanking vowel /u/. Results: Descriptive analysis revealed a strong time-order effect influencing loudness judgments for 7 of the 13 VCV test stimuli. Statistical testing showed no significant didference (P = 0.94) between the relative amplitudes corresponding to the point of equal loudness for plosive-containing versus non-plosive-containing VCV stimuli. However, 6 individual VCV stimuli, containing consonants from 4 separate manners of articulation, produced significant loudness matching data (P < 0.01). Conclusions: The results falsify the hypothesis that plosives, analyzed collectively as a class, possess a heavier perceptual loudness weighting than non-plosive consonants. While 6 individual VCV stimuli indicated potential individual consonantal loudness weightings, these findings must be interpreted cautiously due to the restriction to a single vowel context and the presence of procedural time-order biases.